Back

Genome Biology and Evolution

Oxford University Press (OUP)

Preprints posted in the last 7 days, ranked by how well they match Genome Biology and Evolution's content profile, based on 338 papers previously published here. The average preprint has a 0.18% match score for this journal, so anything above that is already an above-average fit.

1
Utilising nuclear encoded plastid DNA to identify donors of grass-to-grass lateral gene transfer

Bourne, N. G.; Payne, L.; Manzi, S.; Besnard, G.; Vorontsova, M. S.; Jobson, R. W.; Chomicki, G. S.; Dunning, L. T.

2026-08-29 evolutionary biology 10.64898/2026.08.26.747220 medRxiv
Top 0.5%
9.2%
Show abstract

Determining the correct donor species/lineages of grass-to-grass lateral gene transfer (LGT) is vital for deducing specific donor features that could help inform the mechanism of transfer. This requires a dataset spanning a broad range of species to achieve the phylogenetic resolution necessary for precise donor inference. As grass-to-grass LGT often involves the transfer of multi-gene DNA fragments, they can contain additional sequences that allow for accurate orthologous comparisons, such as nuclear DNA of plastid origin (NUPTs). Here we systematically scan for NUPTs in the genomes of four Alloteropsis semialata accessions, whose LGTs have previously been characterised. Using the abundant Panicoideae chloroplast sequences, we reconstruct NUPT phylogenies and infer two lateral acquisitions: one from Paniceae/Digitaria and another from Andropogoneae/Eremochloa adjacent to a previously identified LGT. We then assembled and included an additional 12 Eremochloa chloroplast genomes in the analysis and showed the likely donor was Eremochloa attenuata. Subsequent short-read mapping from E. attenuata to the nuclear region flanking this NUPT showed consistent coverage across the region, including the previously identified LGT, supporting co-transfer. Overall this study highlights the potential for NUPTs to better identify the donors of grass-to-grass LGT.

2
Limited neutral and adaptive genomic divergence suggests Acropora cervicornis can be managed as a single conservation unit across its range

Duffin, P. J.; Ruggeri, M.; Conn, T.; Baums, I. B.; Blanco-Pimentel, M.; Bosch, P.; Carne, L.; Danser, N.; Montoya-Maya, P.; Morikawa, M.; Muller, E. M.; Winters, R. S.; Baker, A. C.; Cunning, R.; Dahlgren, C.; Parkinson, J. E.; Kenkel, C. D.

2026-08-29 genomics 10.64898/2026.08.26.747420 medRxiv
Top 0.5%
8.0%
Show abstract

Genomic signatures can provide key insight into the evolutionary history and remaining adaptive potential of threatened populations. As demographic decline erodes both diversity and the processes maintaining it, understanding how remaining variation is distributed becomes increasingly important for conserving species like the staghorn coral, Acropora cervicornis, a foundational but critically endangered Caribbean reef-builder. We analyzed 46 high-coverage A. cervicornis genomes from 10 locations across the tropical western Atlantic to evaluate neutral and adaptive structure, genomic diversity, demographic history, inbreeding, and connectivity, and generated a regional haplotype reference panel for future genomic monitoring. Genome-wide analyses recovered recurring regional substructure, but differentiation was modest and partly explained by isolation-by-distance and spatial variation in effective migration. Subpopulations had similar levels of genomic diversity, shared demographic history, and limited evidence of local adaptation. These patterns support interpreting sampled Caribbean populations as a single evolutionarily significant unit (ESU) containing multiple regional management units (MUs), rather than as deeply divergent evolutionary lineages. Despite substantial retained variation and low current inbreeding, estimated contemporary effective population size was small, suggesting an increased vulnerability to the effects of drift as demographic collapse continues, especially if structure is reinforced by isolated management. Together, our findings emphasize the urgent need for interventions that preserve and enhance genomic diversity, including risk-managed assisted gene flow. Supported by the haplotype reference panel developed here, these strategies will require coordinated efforts across regional entities to conserve and restore A. cervicornis as a jointly managed, single ESU.

3
Lateral gene transfer shapes the distribution of nitrogen fixation within a cosmopolitan clade of marine Thalassolituus

Barawi, S. S.; LaRoche, J.; Beiko, R. G.

2026-08-29 microbiology 10.64898/2026.08.28.747955 medRxiv
Top 0.8%
5.4%
Show abstract

Biological nitrogen fixation converts dinitrogen gas into ammonia, supplying new bioavailable nitrogen to marine ecosystems, but the evolutionary processes shaping its distribution among heterotrophic bacteria remain unresolved. Thalassolituus, a genus within the family Oceanospirillaceae (order Oceanospirillales), is best known for hydrocarbon degradation, yet nitrogen fixation has been confirmed in only one cultured isolate. We analyzed 74 quality-filtered genomes assigned to Thalassolituus within a broader dataset of 421 Oceanospirillaceae genomes to reconstruct the distribution and evolutionary history of the minimal nifHDKENB gene set. Twenty-five genomes encoded complete or near-complete nif loci and occurred in four well-supported clades interspersed with genomes lacking the pathway. Statistical topology tests rejected the species-tree topology for concatenated NifHDK and NifHDKENB protein alignments, and eleven recombination events across nif loci were supported by at least four detection methods. The core nifHDK gene order remained broadly conserved, but accessory neighborhoods differed among clades, and structural nifHDK genes showed stronger codon adaptation than biosynthesis nifENB genes. Clade 2 combined species-gene tree congruence, conserved gene neighborhoods, and comparatively high nifH codon adaptation, whereas Clades 1 and 4 showed greater phylogenetic discordance, more recombination, and weaker codon adaptation. These results support a reticulate history in Thalassolituus, in which lateral acquisition introduced nitrogen fixation into distinct lineages, vertical inheritance preserved it within some clades, and homologous recombination continued to reshape nif loci. These processes help explain why nitrogen fixation is unevenly distributed among closely related marine heterotrophic bacteria.

4
Quantifying the Rearrangement Complexity of Pangenomes

Bohnenkaemper, L.; Stoye, J.

2026-08-29 bioinformatics 10.64898/2026.08.27.747493 medRxiv
Top 2%
1.4%
Show abstract

The study of evolution between species (phylogenetics) and the study of evolution within a species (population genetics) are highly related, as the same biological mechanisms are fundamental to both fields. Although both have been studied for a long time, their joint study in a unified setting has been prevented by the different time scales they consider and the different data types they employ. A similar discrepancy holds for their whole-genome specializations, comparative genomics and pangenomics. Two active areas in these fields are genome rearrangement studies and graphical pangenomics, respectively. Since the emergence of graphical pangenomics, these have existed as separate fields, despite observations that central data structures representing genomic variants in both fields are highly similar. While there exists a wealth of theoretical results for various rearrangement models in comparative genomics, the application to pangenomic data is hampered by the limitations of rearrangement problem formulations. On the practical side, pangenomes typically contain too many individual genomes for classical problems, such as the often NP-hard parsimony problems, to be solved, or for all-vs-all comparisons using rearrangement distances to be performed. On the theoretical side, some assumptions in the formulation of rearrangement problems, such as the assumption of an underlying tree, are inadequate for many pangenomes. In this work, we propose the Complete Ancestral Reconstruction for Pangenomes (CARP) problem, which overcomes these limitations while retaining intuitive relationships to both classical rearrangement problems and pangenome graphs.

5
A structure-guided classification framework reveals the diversity and catalytic architecture of BECR ribonuclease

Pham, K.; Nicastro, G. G.; Long, A. R.; Aravind, L.; Wilke, C. O.; de Souza, R. F.; Bayer-Santos, E.

2026-08-29 microbiology 10.64898/2026.08.28.747851 medRxiv
Top 3%
1.0%
Show abstract

Microorganisms across all domains of life engage in molecular conflict, deploying toxins to inhibit competitors or respond to biological threats. Among these, ribonuclease toxins are particularly widespread and diverse. A substantial fraction is associated with the BECR fold, a compact /{beta} architecture that supports RNase activity despite extensive divergence. Although several canonical members are well characterized, many BECR-fold proteins remain difficult to identify because of low sequence similarity, variation in catalytic residues, and structural elaborations that obscure evolutionary relationships. The growing availability of high-confidence protein structure predictions provides an opportunity to reassess this deeply divergent protein landscape. Here, we integrate iterative profile-HMM searches, profile-similarity networks, structural analyses, active-site mapping, and genomic context to examine BECR proteins across the tree of life. Our analysis resolves an expanded BECR-fold landscape comprising canonical BECR and BECR-like superfamilies, refines the organization of canonical BECR proteins and identifies previously unrecognized families. We further validate BECR-Tox2 as a toxin neutralized by a cognate immunity protein and show that its homologs occur in both Menshen-like anti-phage systems and polymorphic toxin loci. Together, these findings expand and clarify the BECR-fold landscape and provide a framework for identifying and interpreting highly divergent proteins of this fold.

6
The evolution of family reputation extends indirect reciprocity

Dos Santos, M.; Ohtsuki, H.; Mullon, C.

2026-08-29 evolutionary biology 10.64898/2026.08.27.747476 medRxiv
Top 3%
0.9%
Show abstract

Reputation plays a major role in supporting cooperation among unrelated individuals through indirect reciprocity. By helping others, individuals build a good personal reputation and receive greater benefits from future partners. Most models of indirect reciprocity assume that a person's reputation reflects only their own behaviour. Yet in many societies, people are also judged by their family's reputation. How family reputation affects the evolution of cooperation, and whether reliance on it can itself evolve, remain unclear. Here we show that reputation inheritance expands the conditions under which indirect reciprocity favours cooperation, increasing helping and favouring greater reciprocity. Greater reciprocity in turn favours stronger reliance on inherited reputation, creating a positive feedback that stabilises cooperation, especially when interactions are infrequent or personal behaviour is difficult to observe. This feedback arises because cooperation generates future benefits both for the individual, through their personal reputation, and for their descendants, through inherited reputation. Reputation inheritance thereby provides a route via which kin selection and reciprocity, often treated as alternative explanations for cooperation, can reinforce one another. Our model helps explain why family-based reputation occurs across diverse human societies and provides an evolutionary framework for studying phenomena organised around family standing, including kin-based institutions, feuds between families and honour-based violence within them.

7
Reference-guided comparative genomics of seven Indonesian rice cultivars identifies conserved gene space and trait-associated sequence candidates

Purwestri, Y. A.; Wicaksono, A.; Nurbaiti, S.; Purba, N. T.; Retnaningati, D.; Restiani, R.; Kumalasari, N.; Nuringtyas, T. R.; Handayani, V. D. S.

2026-08-29 genomics 10.64898/2026.08.26.747264 medRxiv
Top 3%
0.8%
Show abstract

Indonesian rice cultivars represent valuable genetic resources, yet many remain poorly characterized at the genomic level. Here, we generated 95.40 Gb of PacBio HiFi sequence data from seven Indonesian rice cultivars and constructed cultivar-specific consensus genomes using the telomere-to-telomere Nipponbare reference AGIS1.0. Sequencing coverage ranged from 27.92x to 41.58x, and the resulting consensus genomes spanned 387.93-390.54 Mb, with BUSCO completeness of approximately 98.3-98.5%. OrthoFinder assigned 99.1% of predicted proteins to 40,737 orthogroups, including 27,514 core orthogroups represented across all seven cultivars, indicating a highly conserved predicted gene space within the reference-guided framework. Targeted analysis recovered 278 of 280 cultivar-by-locus combinations representing 40 genes or gene family entries associated with grain pigmentation, nitrogen and amino-acid metabolism, and starch properties. Comparative predicted protein analysis prioritized ANS1, SBE2b, SSIIa/ALK, Wx/GBSSI, OsAAP6/qPC1, and SSI as candidates for further investigation. Among 269 completed AGIS1.0-anchored promoter comparisons, 159 passed quality-control criteria, whereas 110 were flagged for gene-model, boundary, synteny, or structural concerns. Notably, these flagged comparisons accounted for more than 90% of the alignment-derived sequence variation, emphasizing the importance of rigorous quality control when interpreting apparent promoter divergence. Collectively, these reference-guided genomic resources provide a standardized framework for investigating sequence variation in Indonesian rice germplasm and prioritize testable coding and regulatory candidates for functional validation and future genomics-assisted crop improvement.

8
Can Dental AI Really Beat Dentists? DentalPair-Cert for Rigorous AI-Dentist Inference

Alve, S. R.; Rahman, S.; Meem, S. M. A. C.

2026-09-02 dentistry and oral medicine 10.64898/2026.09.01.26361874 medRxiv
Top 4%
0.5%
Show abstract

A dental AI system and a dentist reading the same radiographs form a paired comparison. Published comparative studies often report the two arms separately against a reference standard, leaving the joint pattern of correctness between them unavailable for secondary paired inference. We show what that omission costs. The accuracy difference remains exactly identified; its sampling variance does not, so the report contains the estimate and not its uncertainty. On a study of 282 units, two published accuracies are consistent with 38 distinct joint tables whose confidence intervals differ in width by a factor of 2.5. The consequence is a three-zone decision map rather than a single threshold: differences at or below 1.06 points are non-significant under every compatible table, differences at or above 6.03 points are significant under every compatible table, and in between the published numbers cannot decide. We then show the omission is repairable at negligible cost. One additional integer, the number of units both arms classify correctly, identifies the joint table exactly and restores standard paired inference. For a panel of readers the pairwise dependences must arise from one joint distribution, a constraint that binds once three readers are present; publishing each reader's joint-correct count against a single reference reader cannot widen and may tighten every pairwise bound, and in a 7-arm experiment reduced them by a median of 37% even for pairs excluding that reference. Where the integer was never published we give DentalPair-Cert, an interval with finite-sample coverage uniformly over every admissible within-unit AI-dentist dependence under the independent-sampling-unit model, certified in both the nuisance maximization and the inversion. Across 4,200,000 simulated comparisons an independence analysis falls to 74.5% coverage with 12.2% type-I error; in a purposive sample of 9 recent comparative studies, 1 reported a paired test on discordant units.

9
PyiTOL: reproducible Python workflows for iTOL annotation and taxonomic monophyly assessment

Zeng, Z.; Wang, Y.

2026-08-29 bioinformatics 10.64898/2026.08.27.747471 medRxiv
Top 4%
0.5%
Show abstract

Motivation: The Interactive Tree of Life (iTOL) is widely used to display and annotate phylogenetic trees, but managing its format-sensitive annotation files impede reproducible high-throughput analyses. Among the maintained Python packages and versions evaluated, none combined template generation, taxonomic monophyly assessment and iTOL batch operations. Results: PyiTOL validates inputs, generates 31 iTOL template schemas (22 accepted by the live batch uploader), performs LCA-based monophyly classification with nested-monophyly detection, sampling-completeness states and polyphyletic subgroup decomposition, plus API upload and session replay. On a topology-constructed benchmark, all calls matched prespecified labels for 4,389 groups; on a 700-genome tree, binary mono/non-mono calls agreed with ETE4 for 409 genera; 17,294 GTDB R232 genera were processed in about 17 s. Availability and Implementation: PyiTOL 1.0.3 (Python [≥]3.10; Linux, macOS and Windows) is MIT-licensed at https://github.com/ZengZichao/PyiTOL and archived with test data at Zenodo (https://doi.org/10.5281/zenodo.22106806).

10
A comprehensive atlas of somatic mutation rates and mutational signatures in normal human cells

Pham, M. H.; Harvey, L. M. R.; Oliver, T. R. W.; Dunstone, E.; Lawson, A. R. J.; Nicola, P. A.; Sanghvi, R.; Hooks, Y.; Mitchell, E.; Jarman, G. L.; Wang, Y.; Abascal, F.; Jung, H.; Neville, M. D. C.; Ishida, Y.; Fowler, J. C.; Le, A. P.; Moody, S.; Marshall, H.; Brzozowska, N.; Ding, C.; Pac, C. A.; Machado, H. E.; O'Neill, L.; Latimer, C.; Humphreys, L.; Saeb-Parsy, K.; Mahbubani, K. T. A.; Baxter, J.; Rassl, D. M.; Vicario, R.; Geissmann, F.; Kabashima, K.; Bleys, R. L. A. W.; Moore, L.; Heer, R.; Coorens, T. H. H.; Behjati, S.; Hoare, M.; Campbell, P. J.; Jones, P. H.; Martincorena, I.; Ra

2026-08-29 genomics 10.64898/2026.08.28.747772 medRxiv
Top 5%
0.3%
Show abstract

Over the course of a lifetime, somatic mutations accrue in normal human cells, causing variation in cell phenotype and engendering somatic evolution with outcomes ranging from the adaptive immune system to cancer. To inform understanding of somatic evolution in the human body we report the mutation rates and mutational signatures of 53 normal cell types. Most show evidence of linear mutation accumulation over time with single base substitution mutation rates ranging from ~3.5/year/diploid genome in spermatogonia and sperm, to ~20/year in postmitotic neurons, ~50/year in mitotically active colorectal epithelial cells, ~60/year in kidney proximal tubule cells and hepatocytes, 100s/year in sun-exposed skin epidermal cells and 10-50/year in the remainder. Certain cell types, including skin epidermis, cardiac myocytes, bladder urothelium, kidney proximal tubule cells, and hepatocytes, show substantial variability in mutation burdens around the linear age trend, indicating the influence of additional factors which differ between individuals and modulate mutation accumulation, including exogenous mutagen exposures. At least 18 single-base substitution and nine small insertion and deletion mutational signatures are present, some in all cell types, some in a subset and others in a single cell type. Known exogenous mutagen exposures and endogenous mutational processes account for some mutational signatures, but the origins and mechanisms underlying many are uncertain. This comprehensive survey of mutagenesis provides a foundation for understanding somatic evolution of human cell populations in health and disease.

11
Genetic dissection of Mycobacteriophage D29 host lysis reveals two lysis regulators and a novel lipoprotein that regulate the lysis event and are localized to distinct regions of the genome

Pollenz, R. S.; Davenport, M.; Ruiz-Houston, K. M.

2026-08-29 microbiology 10.64898/2026.08.27.747656 medRxiv
Top 5%
0.3%
Show abstract

Phage D29 infects Mycobacterium smegmatis mc2 155 and has a non-canonical lysis cassette that encodes two endolysin proteins (Lysin A and Lysin B) and a single two transmembrane domain (TMD) protein, LysA2a similar to F1 cluster phage LysF1a. A 1TMD LysF1b homolog, LysA2b, is encoded by a gene found downstream of the tape measure. Exogenous expression of both LysA2 proteins in tandem is a cytotoxic to M. smegmatis. Deletion of lysA2a produces phages that are lysis competent with a 10-minute triggering delay and 30% plaque size reduction. Deletion of lysA2b results in severe lysis defects manifest by 70% reduced plaque size, delayed lysis timing and reduced burst size. Deletion of both lysA2 genes results in phages that are viable and show lysis phenotypes like the lysF1b deletion. Genetic complementation of lysA2b deleted phage with the lysF1b gene fully complements the lysis phenotypes but alters the triggering time to that of an F1 cluster phage. Energy poisons trigger lysis prematurely in all phages with lysA2 gene deletions. Lysis recovery mutants (LRM) isolated from phages lacking the lysA2b genes generate wild type plaque size and have point mutations that map to TMD1 or the C-terminal region of the lysA2a gene. LRMs isolated from phages lacking both lysA2 genes show premature lysis and have mutations that all map to residue C31 of a novel lipoprotein (gene 64). Deletion of gene 64 does not change wild type D29 lysis phenotypes or rescue the lysis defects of any of the lysA2 mutants. A fitness/competition assay shows that loss of the lysA2 genes imposes a substantial competitive fitness cost. These finding support a lysis regulatory network model where the 2TMD protein is maintained in an inactive state until activated by its cognate 1TMD lysis regulator and the lipoprotein has accessory function that may enhance lysis efficiency.

12
The macroevolutionary impact of an innovation reversal in ray-finned fishes

Brownstein, C.; Harrington, R. C.; Wood, J. E.; Ghezelayagh, A.; Alencar, L.; Munoz, M. M.; Thacker, C. E.; Near, T. J.

2026-08-29 evolutionary biology 10.64898/2026.08.25.747124 medRxiv
Top 5%
0.3%
Show abstract

The evolution of new traits can drive species diversification by facilitating the use of new resources, but environmental change may turn these same adaptations into liabilities.Trait loss is also often associated with the origin of new ecologies, but how losses modulate diversification remains unclear. The swim bladder allows ray-finned fishes to regulate their buoyancy and exploit ecosystems throughout the water column, yet this organ has been lost many times among species-rich lineages. Here, we show that timing and ecological context control the macroevolutionary effects of swim bladder loss. Many lineages of fishes lost the swim bladder over the last 66 million years as they specialized for benthic habitats where buoyancy regulation is unnecessary. Swim bladder loss enabled the descendants of these benthic fishes to diversify in the deep sea where extreme pressure makes its inflation untenable, and in the frigid, oxygen-saturated Southern Ocean, where loss of the oxygen delivery mechanisms required for swim bladder inflation carries little physiological cost. Yet, we detect a selective filter associated with swim bladder loss during extreme global warming 56 to 50 million years ago, when its absence limited the capacity of fishes to escape ecological disruptions on the ocean floor. These contrasting patterns explain how the loss of a complex trait promoted major ecological transitions without increasing overall diversification through deep time. As human activity drives rapid global warming, the evolutionary legacies of swim bladder loss may again shape the fate of marine fish diversity.

13
The QxxR Motif of RNA Helicase Me31B Is Essential for Drosophila Female Fertility and Germline Development

Mansoor, R.; Minhas, A. S.; Thomas, A.; Mansoor, A. A.; McCambridge, A. H.; Dilts, C.; Eshak, J.; Govani, D.; Nylin, B.; Trinidad, J. C.; Kanaan, A. Y.; Kara, E.; Fielder, A.; Fielder, I.; Iglendza, A.; Mukatash, Y.; Pumnea, B.; Menzel, M. M.; Shabazz-Henry, A. L.; Niepielko, M. G.; Gao, M.

2026-08-29 genetics 10.64898/2026.08.27.747641 medRxiv
Top 6%
0.2%
Show abstract

The QxxR motif is evolutionarily conserved within DEAD-box RNA helicases, including Drosophila Me31B and human DDX6, which post-transcriptionally regulate gene expression during animal development. A pathogenic H372R substitution (QxHR to QxRR) in the QxxR motif of human DDX6 has been associated with various developmental defects, but how this motif contributes to DDX6-family protein function remains unclear. Here, we used Drosophila Me31B as an in vivo model to investigate the QxxR motifs developmental role. We generated a Drosophila strain carrying the corresponding H333R missense mutation in Me31B and characterized its effects on female fertility, embryonic viability, germline development, and Me31B-associated molecular pathways. The me31BH333R mutation reduced female fertility in a gene dose-dependent manner, with homozygous mutant females being sterile. Embryos from the mutant females also exhibited primordial germ cell defects. Despite these developmental phenotypes, the me31BH333R mutation did not significantly alter Me31B protein abundance, global ovarian transcriptome or proteome profiles, or representative germ plasm mRNA and protein localization. In contrast, bait-normalized IP-MS analysis revealed altered enrichment of selected Me31B-associated proteins, including increased association of known Me31B interactors Trailer hitch (Tral) and Ypsilon Schachtel (Yps). These findings establish Me31BH333R as an in vivo model for investigating the conserved QxxR motif and suggest that disruption of this motif compromises development not through broad changes in gene expression, but potentially through altered composition or regulation of Me31B-containing ribonucleoprotein complexes.

14
SALRR: Scalable Analysis of Long-Read RNA-Seq Enables Comprehensive Transcriptome Profiling in Human Brain

Kouam, C.; Mingle, J.; Alvarez Jerez, P.; Evans, A.; Moller, A.; Baker, B.; Weller, C.; Paquette, K.; Brooks, J.; Grant, S. M.; Ayuketah, A.; Meredith, M.; Palade, J.; Malik, L.; Hise, K.; Raphael Gibbs, J.; Anderson, J.; Ding, J.; Harbert, R.; Fu, Y.; Zheng, X.; Garcia-Ruiz, S.; Gustavsson, E. K.; Blauwendraat, C.; Ryten, M.; Sedlazeck, F.; Ferrucci, L.; Reed, X.; Nalls, M. A.; Cookson, M. R.; Van Keuren-Jensen, K.; Hutchins, E.; Jain, M.; Billingsley, K. J.

2026-08-29 genomics 10.64898/2026.08.27.747499 medRxiv
Top 6%
0.2%
Show abstract

Isoform-resolved transcriptomics is fundamental to decoding the molecular complexity of the human brain, yet population-scale long-read RNA sequencing has remained inaccessible due to labor-intensive library preparation, sensitivity to RNA degradation in postmortem tissue, and the absence of integrated, reproducible analysis pipelines. Here we present SALRR (Scalable Analysis of Long-Read RNA-seq), an integrated wet-lab and computational platform designed to overcome these barriers. Automated ONT long-read cDNA library preparation on the Hamilton Microlab NGS STAR platform reduces hands-on time by 67% and enables 24 libraries per operator per day while maintaining performance across RNA integrity values. A modular, Snakemake-based pipeline performs end-to-end processing from ONT signal data to isoform-level quantification, incorporating SIRV spike-in calibration, multi-stage quality control, and stringent isoform validation. Applied to 10 postmortem frontal cortex samples from the North American Brain Expression Consortium, SALRR identified 31,607 high-confidence isoforms from 10,075 genes, including 8,532 novel splice variants absent from GENCODE v49, and complex splicing events systematically missed by short-read sequencing at neurodegeneration-relevant loci, including GBA1, CCNF, CHCHD10, and TREM2. All protocols and code are openly available, providing a scalable, community-ready framework for isoform-resolved transcriptomics in neurodegeneration, aging, and complex brain disease.

15
MOSurvivor-Guided Joint CpG Selection and XGBoost Hyperparameter Optimization for Compact Epigenetic Age Prediction

Yelgi, A.; Tavangari, S.; Shakarami, Z.; Janfaza, S.

2026-08-29 genomics 10.64898/2026.08.26.747213 medRxiv
Top 6%
0.1%
Show abstract

Accurate epigenetic age prediction from DNA methylation profiles is intrinsically high-dimensional, creating a need for parsimonious models that preserve predictive performance while reducing the number of assayed cytosine-phosphate-guanine (CpG) loci. This study introduces MOSurvivor, a population-based multi-objective search framework that jointly optimizes a weight-threshold CpG selector and eight XGBoost hyperparameters. Experiments used the GSE40279 whole-blood cohort (656 individuals profiled on the Illumina HumanMethylation450 platform). After retaining 1,000 age-correlated CpGs, five strategies were evaluated on the same 30 seeded 80:20 train/test splits: fixed-parameter XGBoost using all 1,000 CpGs, random search, a genetic algorithm, particle swarm optimization, and MOSurvivor. Internal fitness was estimated using three-fold cross-validation on each training set. Across the 30 held-out test sets, MOSurvivor achieved a mean absolute error (MAE) of 4.149 {+/-} 0.300 years, root mean squared error of 5.545 {+/-} 0.392 years, and R2 of 0.855{+/-} 0.027 while retaining 211.6 {+/-} 54.8 CpGs. Relative to full-feature XGBoost (MAE 4.095 {+/-} 0.285 years), MOSurvivor reduced the feature set by 78.8% at an MAE increase of only 0.054 years (1.3%). Paired Wilcoxon tests found no significant accuracy difference between MOSurvivor and any comparator (all unadjusted p > 0.05; all Holm-adjusted p [≥] 0.476). The most recurrent locus, cg16867657, appeared in 29 runs, whereas mean pairwise Jaccard similarity was 0.124, indicating a small stable core embedded in multiple near-equivalent feature subsets. MOSurvivor thus offers a competitive accuracy-parsimony trade-off rather than superior absolute accuracy. External validation and leakage-free nested feature preselection remain necessary before biological or clinical translation. Keywords: epigenetic clock, DNA methylation, feature selection, multi-objective optimization, XGBoost, metaheuristics, biological aging.

16
Geometric characterization of the HSV - 1 glycoprotein B - amyloid β interaction in Alzheimer's disease using Forman-Ricci curvature

Bou Dagher, L.; Han, Z.; Zhou, S.; Fülöp, T.; Desroches, M.; Rodrigues, S.

2026-08-29 bioinformatics 10.64898/2026.08.26.747308 medRxiv
Top 8%
0.1%
Show abstract

Alzheimer's disease is characterized by the accumulation and aggregation of amyloid-{beta}(A{beta}), but the molecular mechanisms linking environmental and infectious factors to A$\beta$ conformational changes remain incompletely understood. Herpes simplex virus type 1 (HSV-1) has been proposed as a potential contributor to AD pathology, and interactions between the viral glycoprotein B (gB) and A$\beta$ may influence the conformational behaviour of the peptide. Molecular dynamics (MD) simulations provide atomic-scale information on such interactions, but conventional structural descriptors may not fully capture changes in the organization of residue interaction networks. Here, we introduce a graph-geometric framework based on Forman-Ricci curvature to characterize the evolution of residue interaction networks during MD simulations. Each simulation frame is represented as a residue interaction graph based on C--C contacts, and residue-wise curvature profiles are analysed across time. We apply the framework to A{beta}1-42 in isolation and in complex with HSV-1 gB. Conventional MD analyses indicate stable association of the simulated complex, favourable interaction energetics, and conformational changes in A{beta}, including a transition from -helical structure toward {beta}-turn-rich conformations over the simulated timescale. Forman-Ricci curvature reveals pronounced and spatially localized remodelling of the A{beta} residue interaction network in the complex, with the strongest changes concentrated in the C-terminal region. These regions also exhibit reduced temporal curvature fluctuations and progressively distinct geometric behaviour throughout the simulation. Hierarchical clustering further identifies cooperative groups of residues with coordinated curvature dynamics, including a prominent C-terminal domain. Together, these results demonstrate that Forman-Ricci curvature provides a complementary description of biomolecular dynamics by capturing changes in the geometric organization of residue interaction networks that are not directly represented by conventional structural descriptors. The framework provides a general computational approach for studying network-level structural remodelling in protein molecular dynamics and offers a quantitative perspective on the conformational consequences of HSV-1 gB--A{beta} association.

17
Bacterial metagenome in plaque, saliva, and tumor samples from individuals with and without OSCC by next-generation sequencing

ERIRA, A.; ROBAYO, D. A. G.; GAMBOA, F.; CHALA, A.; MORENO, A.; ARREGUI, A. C.; MUNOZ, E.; NOGUERA, J.; TOBAR-TOSSE, F.

2026-08-29 bioinformatics 10.64898/2026.08.27.747557 medRxiv
Top 8%
0.1%
Show abstract

Background: Oral dysbiosis has been associated with oral squamous cell carcinoma (OSCC); however, most microbiome studies rely on 16S ribosomal RNA (rRNA) gene sequencing, limiting species-level taxonomic resolution. Methods: Dental plaque, saliva, and tumor tissue samples from 10 patients with OSCC and dental plaque and saliva samples from 10 healthy controls were analyzed in this exploratory cross-sectional study. DNA was extracted and subjected to shotgun metagenomic sequencing using the Illumina MiSeq platform. Sequence reads were quality filtered with fastp, taxonomically classified using Kraken2 v2.1.3, and species-level abundances were re-estimated with Bracken v2.9 following the removal of human reads and low abundance taxa. Relative abundances were compared using the Mann Whitney U test with the Benjamini Hochberg false discovery rate correction, while the Bray Curtis principal coordinate analysis was used as an exploratory approach to visualize microbial community patterns. Results: Shotgun metagenomic sequencing revealed distinct bacterial community profiles across the oral microenvironment. Dental plaque exhibited the highest taxonomic diversity and relative abundance. The control plaque was enriched in Streptococcus koreensis, Capnocytophaga sp. oral taxon 878, Treponema sp. Marseille Q4132, and Leptotrichia sp. oral taxon 498, whereas the plaque from patients with OSCC showed a higher relative abundance of Pyramidobacter piscolens, Parvimonas parva, and Gemella sanguinis. Salivary samples displayed lower diversity and a more homogeneous composition, predominantly comprising Capnocytophaga endodontalis, Prevotella jejuni, Aggregatibacter aphrophilus, and Gemella sanguinis. The tumor tissue showed relatively higher abundance of Sellimonas catena, Escherichia coli, Solobacterium moorei, and Lacrimispora sp. HJ 01. Conclusions: This exploratory study provides species-level characterization of the oral microbiome across multiple oral microenvironments in OSCC and generates hypotheses for future integrative metagenomic and functional studies investigating the potential contribution of oral bacterial communities to OSCC pathogenesis.

18
Perturb-seq identifies co-regulated gene programs shaping hematopoietic stem and progenitor cell function

Bowness, J. S.; Bernal Martinez, A.; Barinka, J.; Schulte-Schrepping, J.; Renders, S.; Waclawiczek, A.; Leppa, A.-M.; Trumpp, A.; Raffel, S.; Haas, S.; Velten, L.

2026-08-29 genomics 10.64898/2026.08.27.747033 medRxiv
Top 9%
0.1%
Show abstract

To sustain blood formation, hematopoietic stem and progenitor cells (HSPCs) coordinate a multitude of cell biological processes, from cell cycle control and stress responses to lineage priming. While many genetic regulators of high-level HSPC function have been identified, how HSPCs coordinate more basal cell biological programs, and how such programs relate to stem cell function, remains incompletely understood. Here we use Perturb-seq to profile the transcriptional consequences of targeting 520 genes by CRISPRi in primary mouse HSPC cultures. We developed an analytical strategy to separate perturbation-induced changes in cell-state abundance and clonal heterogeneity from cell-state-local transcriptional effects. From these local perturbation signatures, we identified 19 gene regulatory programs (GRPs) that are defined by co-regulation in response to genetic perturbation, in contrast to co-expression or human curation, and align well with cell biological processes. By decomposing gene expression data from functional and clinical studies into program activity, we show that GRP activities associate with, and predict, phenotypes such as clonal output after transplantation, as well as survival and drug response in retrospective acute myeloid leukemia (AML) cohorts. Together, our study establishes perturbation-derived co-regulation programs as an interpretable framework for linking genetic regulators, cell-biological processes and stem-cell-associated phenotypes.

19
Unravelling genomic and functional traits of two biocontrol and plant growth-promoting Pseudomonas endophytes

Santoyo, G.; Flores, A.; Castelan-Sanchez, H. G.; Valenzuela-Ruiz, V.; de los Santos-Villalobos, S.; Mitra, D.; Babalola, O. O.; Schoebitz, M.; Orozco-Mosqueda, M. d. C.

2026-08-29 microbiology 10.64898/2026.08.28.747936 medRxiv
Top 9%
0.1%
Show abstract

Plant growth-promoting bacterial endophytes represent a sustainable strategy for enhancing agricultural productivity while reducing reliance on synthetic fertilizers and pesticides. This study focused on the genomic and functional characterization of two endophytic bacterial strains, R11F and R19M, isolated from bean and maize roots, respectively. Comparative analyses based on 16S rRNA gene sequences, average nucleotide identity (ANI), and genome-to-genome distance calculations (GGDC) classified both isolates as Pseudomonas palleroniana. Comparative genomic analyses revealed highly conserved genomes containing genes associated with plant colonization, phosphate solubilization, stress adaptation, heavy metal resistance, and hydrocarbon degradation. Genome mining further identified 17 and 18 biosynthetic gene clusters (BGCs) in R11F and R19M, respectively, including non-ribosomal peptide synthetases (NRPS), pyoverdine, NRP-metallophores, RiPP-like compounds, arylpolyenes, {beta}-lactones, terpenes, NAGGN, and hydrogen cyanide. Strain-specific BGCs associated with syringomycin and viscosin biosynthesis were identified in R11F, whereas R19M harbored clusters related to asplenin and kolossin biosynthesis. In vitro assays confirmed indole production, phosphate solubilization, and siderophore production, as well as the ability of both strains to grow in nitrogen-free medium. Both strains significantly inhibited the growth of Fusarium oxysporum, Phytophthora cinnamomi, and Colletotrichum gloeosporioides. Furthermore, plant inoculation assays demonstrated host-dependent growth promotion, with R11F showing the most consistent improvements in plant growth parameters in tomato, wheat, and lentil. Overall, the integration of comparative genomics and experimental validation demonstrates that P. palleroniana R11F and R19M possess complementary traits associated with plant growth promotion, pathogen suppression, saline stress adaptation, and bioremediation.

20
A Simple Method to Distinguish Active and Inactive Aptamers by Analyzing the Ruggedness of the Aptamer Free Energy Landscape

Subramanian, G.; Thiel, W.; Singh, R.

2026-08-29 bioinformatics 10.64898/2026.08.26.747184 medRxiv
Top 10%
0.1%
Show abstract

Aptamers are structured nucleic acid ligands capable of high affinity, high specificity molecular recognition generated using variations of the SELEX (Systematic Evolution of Ligands by Exponential Enrichment) process. However, SELEX often produces sequences that enrich yet may lack binding efficacy. We propose a measure called the Ruggedness Composite Index (RCI) along with a method for computing it, that can be used to distinguish binding-competent ('active') aptamers from weak or non-binding ('inactive') aptamers. Given a set of aptamers, RCI incorporates information on their fragmentation (landscape partitioning), basin entropy (metastable state distribution), cumulative density irregularity (non-uniform occupancy), and structural energy correlation length (structure-energy coupling scale). We test whether secondary-structure folding energy landscape topology distinguishes active from inactive aptamers using a multiscale level set framework across six datasets. Active aptamers show lower RCI values and occupy smoother, funnel-like conformational spaces, while inactive aptamers show higher RCI values, reflecting fragmented, high-entropy landscapes. By contrast, classical thermodynamic features, such as minimum free energy, show limited discrimination between active and inactive aptamers. In all datasets, sequences that exhibit enrichment which is not monotonic but lack specificity exhibit elevated ruggedness, indicating landscape topology can predict non-specific enrichment. These results indicate that folding landscape organization can be used as a predictor of aptamer activity and establish RCI as a simple, mechanistically interpretable measure for improving candidate prioritization, especially in therapeutic aptamer discovery.